Papers with visual-semantic embedding space
Multi-Head Attention with Diversity for Learning Grounded Multilingual Multimodal Representations (D19-1)
Copied to clipboard
| Challenge: | Recent studies have advanced learning VSE under the monolingual setup. |
| Approach: | They propose a model with diverse multi-head attention to learn grounded multilingual multimodal representations by leveraging visual object detection. |
| Outcome: | The proposed model performs well in German-Image and English-Image matching tasks and in the Semantic Textual Similarity task with English descriptions of visual content. |